Papers by Hellina Hailu Nigatu
Evaluating Machine Translation Datasets for Low-Web Data Languages: A Gendered Lens (2026.findings-acl)
Copied to clipboard
Hellina Hailu Nigatu, Bethelhem Yemane Mamo, Bontu Fufa Balcha, Debora Taye Tesfaye, Elbethel Daniel Zewdie, Ikram Behiru Nesiru, Jitu Ewnetu Hailu, Senait Mengesha Yayo
| Challenge: | afan oromo, amharic, and tigrinya are low-resourced languages . they are used for training, benchmarks, news, health, and sports . afono o'mara: quantity does not guarantee quality of MT datasets . |
| Approach: | They investigate the quality of machine translation datasets for three low-resourced languages . they found a large skew towards the male gender in the datasets . |
| Outcome: | The results show that training data has large representation of political and religious text, but benchmark datasets focus on news, health, and sports. |
A Case Against Implicit Standards: Homophone Normalization in Machine Translation for Languages that use the Ge’ez Script. (2025.emnlp-main)
Copied to clipboard
Hellina Hailu Nigatu, Atnafu Lambebo Tonja, Henok Biadglign Ademtew, Hizkiel Mitiku Alemayehu, Negasi Haile Abadi, Tadesse Destaw Belay, Seid Muhie Yimam
| Challenge: | Homophone normalization is a pre-processing step used in Amharic natural language processing (NLP) but it also results in models that are unable to process different forms of writing in a single language. |
| Approach: | They propose a method where normalization is applied to model predictions instead of training data and a scheme where normalized data is preserved in training. |
| Outcome: | The proposed model achieves an increase in BLEU score of up to 1.03 while preserving language features in training. |
Cognate Detection for Historical Language Reconstruction of Proto-Sabean Languages: the Case of Ge’ez, Tigrinya, and Amharic (2025.coling-main)
Copied to clipboard
| Challenge: | As languages evolve, we risk losing ancestral languages. |
| Approach: | They propose to use cognates to reconstruct proto-languages from cognates in child languages that have likely evolved from the same word in the proto-linguistics. |
| Outcome: | The proposed method is based on automatic cognate detection and in-context learning with GPT-4o to generate the proto-language from the cognates and use Sequence-to-Sequence models. |
The Zeno’s Paradox of ‘Low-Resource’ Languages (2024.emnlp-main)
Copied to clipboard
| Challenge: | 'low resource' languages are understudied by the NLP community, while 'high resource' is referred to as 'achieved', while high-resource languages are referred . |
| Approach: | They qualitatively analyzed 150 papers from the ACL Anthology and popular speech-processing conferences that mention the keyword ‘low-resource. |
| Outcome: | The proposed analysis reveals that several interacting axes contribute to ‘low-resourceness’ of a language and why that makes it difficult to track progress for each individual language. |
mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languages (2025.findings-acl)
Copied to clipboard
| Challenge: | Knowledge Graphs are structured multirelational graphs that store factual knowledge. |
| Approach: | They introduce a Retrieval-Augmented Generation (mRAKL) based system to perform mKGC. |
| Outcome: | The proposed approach improves over a no-context setting with an idealized retrieval system. |
Viability of Machine Translation for Healthcare in Low-Resourced Languages (2025.emnlp-main)
Copied to clipboard
Hellina Hailu Nigatu, Nikita Mehandru, Negasi Haile Abadi, Blen Gebremeskel, Ahmed Alaa, Monojit Choudhury
| Challenge: | MT errors are more pronounced in low-resourced languages where human translators are scarce and MT tools perform poorly. |
| Approach: | They propose to use a publicly available machine translation system to analyze machine translation errors in healthcare domains. |
| Outcome: | The proposed system reduces errors in two low-resourced languages for healthcare. |